Skip to content

Describe running, upgrading and removing the Docker image in the Installation Guide - #1179

Open
vharseko wants to merge 3 commits into
OpenIdentityPlatform:masterfrom
vharseko:docs/docker-in-asciidoc
Open

vharseko wants to merge 3 commits into
OpenIdentityPlatform:masterfrom
vharseko:docs/docker-in-asciidoc

Conversation

@vharseko

@vharseko vharseko commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

Problem

The AsciiDoc guides say almost nothing about the Docker image. Docker comes up once, in passing, in install-guide/chap-upgrade.adoc ("as do the Docker image and the native packages, which run upgrade --no-prompt --force"). The Installation Guide has procedures for ZIP, GUI, CLI, .deb, .rpm and MSI, but none for Docker: not for installing, not for upgrading, not for removing. The replication and certificates chapters do not mention that the image manages both itself.

Change

Each new piece links to opendj-packages/opendj-docker/README.md. The README stays the only place that describes the environment variables, the health check, replication and certificates of the image, so this PR does not repeat any of that.

Installation Guide

  • chap-install.adoc: a new procedure, To Run OpenDJ in Docker (#install-docker). It covers:

    • the image names and tags;
    • docker run with a volume at /opt/opendj/data and --init, and why the volume is needed: the image declares no VOLUME;
    • ports published on loopback only, with a warning to set your own ROOT_PASSWORD on the first start before you publish them on other interfaces: the image reads it only when it creates the instance, except for a container replicated with OPENDJ_REPLICATION_TYPE=simple, whose join binds with it on every start;
    • waiting for healthy (or unhealthy once the 5-minute start period is over), then a check with ldapsearch;
    • after a failed first start, removing the container and its volume before trying again: a restart over that volume can report healthy without the backend (Docker image: a restart after a failed first start reports healthy over a half-bootstrapped volume #1182).

    "To Prepare For Installation" also points to it.

  • chap-upgrade.adoc: a new procedure, To Upgrade a Docker Container (#upgrade-docker). Its steps:

    1. Check that the instance is on a volume, and move it to one with docker cp … - | tar -x if it is not. The following steps use the volume name that this check prints; a bind mount is backed up as a host directory, after the container is stopped.
    2. Back up the volume with docker stop -t 60 and a tar under umask 077. The archive holds keys and password hashes.
    3. Start a new container with all the options of the old one.
    4. For an instance with reject-unauthenticated-requests:true, set HEALTHCHECK_BIND_DN/HEALTHCHECK_BIND_PASSWORD_FILE. Images before 5.2.0 probed as the root user, the current one probes anonymously ([#1092] Probe the Docker container's health without binding as the root user #1102).
    5. Wait for healthy, and expect long upgrade tasks to keep the container unhealthy past the start period. After a failed upgrade, do not restart the container: a failed post-upgrade task, such as an index rebuild, leaves the new version recorded, so the next start reports healthy (Docker image: a restart after a failed post-upgrade task reports healthy, with the upgrade never completed #1185). Revert, or rebuild the indexes named in the logs; upgrade.log has the details.

    It also gives a revert, which restores into an empty volume, and points to #upgrade-repl for rolling upgrades. The passing mention of the Docker image now links to this procedure.

  • chap-uninstall.adoc: a new procedure, To Remove a Docker Container (#uninstall-docker). It says when dsreplication disable is still needed: only OPENDJ_REPLICATION_TYPE=simple with REPLICATION_PEERS removes a server registered by name. It gives the order: remove the container first, then recreate every other container without it in REPLICATION_PEERS. The environment of a container cannot be changed, so a restart is not enough. The last step removes the volume.

  • preface.adoc: one sentence pointing "just trying it" readers to the Docker procedure.

Administration Guide

  • chap-replication.adoc: a NOTE saying that the image joins the topology itself and that, with REPLICATION_PEERS, it removes servers that are not listed. It says how to add a server: list it in every container first, then start it. It says how to remove one: remove its container first, then update the list everywhere. A server registered by an address is never removed.
  • chap-change-certs.adoc: a NOTE on the image's certificate handling:
    • Only the key*/trust* files on SECRET_VOLUME are copied into the instance, on start and, by default, while the server runs.
    • admin-keystore, admin-truststore and ads-truststore stay in the instance and are replaced with the procedures of this chapter.
    • A renewed keystore is served without a restart only if it keeps the previous alias.

Testing

  • All six chapters render with AsciidoctorJ 2.5.3 (-v) without warnings. The new anchors resolve, {opendj-version} is substituted, and {{.State.Health.Status}} stays verbatim.
  • Ran on the released openidentityplatform/opendj image (5.1.2): the install docker run, the ldapsearch check (it prints dn: dc=example,dc=com), the volume backup, a new container of the same image over the same volume (its start ran upgrade, which is a no-op on equal versions, and the data was kept), and the removal steps.
  • Ran on stand-in alpine containers whose files are owned by 1001:0: the docker inspect mount check, the backup under umask 077 (the archive is -rw-------), the restore into a fresh volume, and the move with docker cp … -. Owner and modes are kept, including on the volume root.
  • Ran on an image built from master 3f4deb9178: after a first start that failed in create-backend, docker restart reported healthy with only adminRoot, which is why the install procedure says to remove the volume (Docker image: a restart after a failed first start reports healthy over a half-bootstrapped volume #1182).
  • Not checked on an image: an upgrade from one version to another, which is the road To Upgrade a Docker Container documents. No CI step runs it either (Docker image: no CI step starts the build image over a volume written by a released image #1183). Nor a failed post-upgrade task followed by a restart: the advice not to restart after a failed upgrade rests on Upgrade.java (changeBuildInfoVersion at :877 comes before the post-upgrade tasks at :887-891) and on run.sh:142/:161 (Docker image: a restart after a failed post-upgrade task reports healthy, with the upgrade never completed #1185).
  • Not checked on an image: that healthy comes only after the whole bootstrap. The 5.1.2 image still has the old health check, which turned healthy before import-ldif had finished. The text follows healthcheck.sh on master, which is what the next release ships.

@vharseko vharseko added enhancement docker docs upgrade Upgrading between versions and migrating from other directory servers labels Oct 8, 2026
@vharseko
vharseko requested a review from maximthomas October 8, 2026 07:47

@maximthomas maximthomas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise: The upgrade and revert procedures handle the parts that are easy to get wrong.

  • The backup runs under umask 077 in a throwaway container and says why: the archive holds the server's keys and the password hashes (chap-upgrade.adoc:331-338).
  • The revert restores into an empty volume (docker volume rm opendj-data before the extract) and gives the reason: files that the upgrade added would otherwise stay behind (chap-upgrade.adoc:366-374).
  • The upgrade step warns that from 5.2.0 the health check searches anonymously, so an instance with reject-unauthenticated-requests:true needs HEALTHCHECK_BIND_DN/HEALTHCHECK_BIND_PASSWORD_FILE (chap-upgrade.adoc:354).

issue (non-blocking): The guide presents ROOT_PASSWORD as the root password, but the image reads it on the first start only.

opendj-doc-generated-ref/src/main/asciidoc/install-guide/chap-install.adoc:679

run.sh reads ROOT_PASSWORD only on the bootstrap road (run.sh:183, bootstrap/setup.sh:51). A container started over an existing instance goes through run.sh:137-174 (upgrade, marker, start_server) and never reads it. Take a reader who ran the example with password and then follows ":679 before you publish them on other interfaces, set a ROOT_PASSWORD of your own": they recreate the container with a new ROOT_PASSWORD and the ports on all interfaces. The server comes up healthy and exposed, and its root password is still password. The README says "only the initial root password" (README.md:34); the guide does not.

The example publishes these ports on the loopback interface of the host only; before you publish them on other interfaces, set a `ROOT_PASSWORD` of your own on the first start. The image reads `ROOT_PASSWORD` only when it creates the instance: a container started over an instance that is already there keeps the root password of that instance, which you change with `ldappasswordmodify` as on any other server. The root user is `cn=Directory Manager`, with the initial password `ROOT_PASSWORD`, and the base DN is `dc=example,dc=com`; without `ADD_BASE_ENTRY` that suffix is created empty.

suggestion (non-blocking): The guide does not say what to do after a failed first start. A later start over the same volume can report healthy on a half-bootstrapped instance.

opendj-doc-generated-ref/src/main/asciidoc/install-guide/chap-install.adoc:689, :702

Suppose setup has written data/config and then dsconfig create-backend (bootstrap/setup.sh:86-88) or the import fails. The first start stays unhealthy (run.sh:195-198), as the text says. The next start can come from docker start, which :702 offers, from docker restart, from a restart policy, or from docker rm plus docker run with fixed variables. In every case it deletes the marker (run.sh:35), finds ./data/config (:137) and runs upgrade -n --force. On equal versions that exits 0 (Upgrade.java:994-999), and :161 writes the marker. The container then reports healthy with no userRoot backend or no base entry. That contradicts ":689 healthy once all of this has succeeded".

A first start that fails never reports healthy: what failed is in `docker logs opendj`. To try again, remove the container and its volume, with `docker rm -f opendj` and `docker volume rm opendj-data`, then start over: a container started over the volume of a failed first start takes it for an installed instance, and can report itself healthy without the backend or the base entry.

suggestion (non-blocking): The backup, step 3 and the revert hard-code the volume opendj-data. Nothing tells the reader to use the name step 1 printed, and a bind mount cannot be backed up this way.

opendj-doc-generated-ref/src/main/asciidoc/install-guide/chap-upgrade.adoc:313-319, :337, :350, :372-374

Step 1 lists {{.Name}} {{.Destination}}. It counts the instance as on a volume when any mount is at /opt/opendj/data, so a compose-prefixed or anonymous volume passes, and so does a bind mount, which prints an empty name. Step 2 then runs -v opendj-data:/data as written. Docker creates a new empty opendj-data, tar archives only ./ and exits 0, and the revert later restores that empty archive. Step 3's "over the same volume, with every other option of the old container" covers step 3 only.

The following steps use the volume `opendj-data`: use the name that the command printed instead. An empty name is a bind mount of a host directory, which the following command shows; back up that directory instead of running the commands of the next step, and mount it again in step 3:
+

[source, console]
----
$ docker inspect -f '{{range .Mounts}}{{.Source}} {{.Destination}}{{println}}{{end}}' opendj
----

suggestion (non-blocking): No CI step starts the build image over a volume written by a released image, so the road this procedure documents has no observable.

opendj-doc-generated-ref/src/main/asciidoc/install-guide/chap-upgrade.adoc:306

CI reuses volumes only between containers of the same build image (build.yml:640/:675, docker-test-replication.sh). Released images run only in the benchmark (build.yml:847/:1222), with no volume. A restart on the same image does reach run.sh:142, but Upgrade.isVersionCanBeUpdated returns success on equal versions before any task runs (Upgrade.java:994-999). A 5.1.x to 5.2.0 regression in the image's upgrade-on-start therefore stays green in every cell. The PR's own run, 5.1.2 over 5.1.2, was the same no-op.

      - name: Docker test upgrade from a released image
        shell: bash
        run: |
          docker run -d --name test_upgrade --init -e ADD_BASE_ENTRY=--addBaseEntry \
            -v test_upgrade_data:/opt/opendj/data openidentityplatform/opendj:5.1.2
          timeout 5m bash -c 'until docker exec test_upgrade /opt/opendj/bin/ldapsearch --port 1389 --bindDN "cn=Directory Manager" --bindPassword password --baseDN dc=example,dc=com --searchScope base "(objectClass=*)" dn 2>/dev/null | grep -q "^dn: dc=example,dc=com"; do sleep 5; done'
          docker stop -t 60 test_upgrade && docker rm test_upgrade
          docker run -d --name test_upgrade --init -v test_upgrade_data:/opt/opendj/data "$IMAGE"
          timeout 10m bash -c 'until [ "$(docker inspect -f "{{.State.Health.Status}}" test_upgrade)" = healthy ]; do sleep 5; done'
          docker exec test_upgrade /opt/opendj/bin/ldapsearch --port 1389 --bindDN "cn=Directory Manager" --bindPassword password \
            --baseDN dc=example,dc=com --searchScope base "(objectClass=*)" dn | grep -q "^dn: dc=example,dc=com"

Pin: the step goes red when the new image fails to upgrade or serve a 5.1.2 volume. Not run. The first wait polls the base entry because the 5.1.2 health check turned healthy before import-ldif ended.


nitpick (non-blocking): "so only the administrators of the server may be able to read it" reads as a possibility, not as a requirement.

opendj-doc-generated-ref/src/main/asciidoc/install-guide/chap-upgrade.adoc:331

The archive holds the private keys of the server and the password hashes of its users, so make sure that only the administrators of the server can read it:

@vharseko

vharseko commented Oct 8, 2026

Copy link
Copy Markdown
Member Author

Thank you for the review. All five points hold against the code. Four of them are in f5e8ca0; the CI step is tracked separately.

ROOT_PASSWORD on the first start only. Confirmed: run.sh:183 comes after the branch for an existing instance, which ends in exec start_server. Taken as suggested, with one more sentence: a replicated container binds with ROOT_PASSWORD to join its topology on every start. That sentence links to the Replication section of the README instead of repeating what the README says about changing the root password there.

Restart after a failed first start. Confirmed, and reproduced on an image built from master: with BACKEND_TYPE=bogus the first start fails in create-backend. After docker restart the container reports healthy, list-backends shows adminRoot only, and a search of the base DN returns 32. Taken as suggested. The guide can only describe how to work around this; the image itself should not report such a volume healthy, which is now #1182.

Hard-coded volume name. Confirmed: an empty {{.Name}} is a bind mount. In that case -v opendj-data:/data archives a new, empty volume and exits 0. Taken as suggested. The sentence that follows now starts "If nothing is mounted at /opt/opendj/data", so that it no longer reads as a reply to the new paragraph.

No CI for an upgrade from a released image. Agreed. You are also right that my own run, 5.1.2 over 5.1.2, upgraded nothing; I have corrected the Testing section of the description. The step tests the image rather than this text, so it is #1183, with your sketch. If it finds a 5.1.x volume that the current image cannot upgrade, that deserves its own fix rather than holding this documentation back.

"may be able to read". Taken as suggested.

@maximthomas maximthomas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise: The round-1 points landed where the problems were, and the image-side work went to its own issues.

  • chap-upgrade.adoc:319-325 now names the host directory of a bind mount ({{.Source}}) instead of letting -v opendj-data:/data archive a new, empty volume.
  • The failed-first-start remedy at chap-install.adoc:689 comes with a reproduction (BACKEND_TYPE=bogus, list-backends showing adminRoot only), and the image fix is tracked as #1182.
  • "If nothing is mounted at /opt/opendj/data" (chap-upgrade.adoc:327) keeps the copy branch from reading as a reply to the new paragraph.

issue (non-blocking): The bind-mount sentence in upgrade step 1 tells the reader to skip all of step 2's commands, and docker stop is one of them.

opendj-doc-generated-ref/src/main/asciidoc/install-guide/chap-upgrade.adoc:319, :344

Taken literally, a reader with a bind mount backs up the host directory while the server is still writing to it, then skips the stop. The only command that fails on the running container is docker rm opendj in step 3 (:354), and it runs after the copy is taken. That copy is all the revert (:374) can restore. Step 2's title and :339 ("the files of the stopped instance") say to stop first, so the text contradicts itself, and the literal reading is the unsafe one. This is the wording I suggested in round 1; the mistake is mine.

The following steps use the volume `opendj-data`: use the name that the command printed instead. An empty name is a bind mount of a host directory, which the following command shows. In the next step, stop the container with its first command, then back up that directory instead of running the second one, and mount the directory again in step 3:

issue (non-blocking): "A container whose upgrade fails never reports healthy" stops being true if a post-upgrade task failed and the container is started again.

opendj-doc-generated-ref/src/main/asciidoc/install-guide/chap-upgrade.adoc:372; opendj-server-legacy/src/main/java/org/opends/server/tools/upgrade/Upgrade.java:816, :877, :887-890, :991-999; opendj-packages/opendj-docker/run.sh:142

Upgrade.upgrade records the new version (changeBuildInfoVersion, :877) before it runs the post-upgrade tasks (:887-890). If one of them fails (an index rebuild, or the DN equality verify), upgrade exits 1 and run.sh starts the server without the marker, so the container is unhealthy, as the text says. The next start over that volume (docker restart, a restart policy, or docker rm + docker run) runs upgrade -n --force again. isVersionCanBeUpdated (:816, :991-999) then finds the versions equal and exits 0, run.sh writes the marker, and the container reports healthy with those indexes still as the failure left them. On the image side this is the upgrade sibling of #1182. The guide can at least say not to restart past the failure.

A container whose upgrade fails does not report itself healthy, and `docker logs opendj` shows why. Do not start it again to get past the failure: when an index rebuild that runs after the upgrade fails, the new version is already recorded, so the next start skips the upgrade and reports healthy with those indexes still untrusted. Revert as follows, or rebuild the indexes that `docker logs opendj` names before you rely on the container.

issue (non-blocking): The two new ROOT_PASSWORD sentences contradict each other, and the second one is wrong for srs, sdsr and rg.

opendj-doc-generated-ref/src/main/asciidoc/install-guide/chap-install.adoc:679; opendj-packages/opendj-docker/run.sh:60-61, :169, :203

With OPENDJ_REPLICATION_TYPE=simple and REPLICATION_PEERS or MASTER_SERVER set, run.sh:169 starts join.sh on every start of an installed instance, and join.sh binds with ROOT_PASSWORD (join.sh:79, :150). So "only when it creates the instance" does not hold for that container. The one-shot types run replicate.sh only in the bootstrap branch (run.sh:203, and the README paragraph on one-shot types), so "on every start" does not hold for them. Both errors lean towards caution. The first sentence is my round-1 wording.

Apart from the replication join described next, the image reads `ROOT_PASSWORD` only when it creates the instance: [...] A container replicated with `OPENDJ_REPLICATION_TYPE=simple` binds with `ROOT_PASSWORD` to join its topology on every start, so read [...] before you change the root password of one.

nitpick (non-blocking): "docker inspect reports starting until then" leaves out unhealthy.

opendj-doc-generated-ref/src/main/asciidoc/install-guide/chap-install.adoc:681; opendj-packages/opendj-docker/Dockerfile:93

With --start-period=5m --interval=30s --retries=3, a first start that fails, or that runs longer than the start period, turns unhealthy about six minutes in. The timeout 5m loop gives up silently before that.

. Wait until the container reports itself healthy; `docker inspect` reports `starting` until then, and `unhealthy` once the 5-minute start period is over while the checks still fail:

nitpick (non-blocking): The remedy says the next container "takes" the volume of a failed first start for an installed instance, but that happens only once setup has written data/config.

opendj-doc-generated-ref/src/main/asciidoc/install-guide/chap-install.adoc:689; opendj-packages/opendj-docker/run.sh:137; opendj-packages/opendj-docker/bootstrap/setup.sh:31-34

run.sh:137 takes the installed-instance branch when [ -d ./data/config ] is true. A first start that fails before /opt/opendj/setup runs (setup.sh:31-34 refuses a ROOT_PASSWORD with a line break) bootstraps again on the next start. The remedy is right either way; only the clause overstates. This is my round-1 wording again.

[...] a container started over the volume of a failed first start can take it for an installed instance and report itself healthy without the backend or the base entry.

…allation Guide

Add the procedures "To Run OpenDJ in Docker", "To Upgrade a Docker Container"
and "To Remove a Docker Container", point the preface to the first of them,
and add notes to the replication and certificates chapters of the
Administration Guide about what the image manages itself. The environment
variables, health check, replication and certificates of the image stay
described in opendj-packages/opendj-docker/README.md only, which the new
text links to instead of repeating it.
…o after a failed first start, and which volume the Docker upgrade backs up

The Docker image reads ROOT_PASSWORD only when it creates the instance,
and a container started over the volume of a failed first start can
report itself healthy without the backend (OpenIdentityPlatform#1182), so the install
procedure says to remove that volume before trying again. The upgrade
procedure uses the volume name that its first step prints, and backs up
a bind mount as a host directory instead of archiving a fresh empty
volume.
…restart a container past a failed upgrade, and to stop the container before backing up a bind mount
@vharseko
vharseko force-pushed the docs/docker-in-asciidoc branch from f5e8ca0 to 6510d4e Compare October 9, 2026 07:32
@vharseko

vharseko commented Oct 9, 2026

Copy link
Copy Markdown
Member Author

Thank you for the second round. All five points hold against the code; all are in 6510d4e. The branch is also rebased onto the current master.

Bind mount skips docker stop. Confirmed. Taken as suggested.

Restart past a failed upgrade. Confirmed: changeBuildInfoVersion (Upgrade.java:877) runs before the post-upgrade tasks, a failed rebuild exits 1, and the next upgrade stops at isVersionCanBeUpdated with success. Taken with two changes. The text says "as the failed rebuild left them" instead of "still untrusted", since a rebuild that fails before it starts leaves them trusted. It also points to /opt/opendj/data/logs/upgrade.log, since docker logs only says "Please check log for further details". The image side is #1185. #1184 does not cover it, because its BOOTSTRAP_PENDING is written by the bootstrap only.

ROOT_PASSWORD sentences. Confirmed, and the second sentence was mine. Taken as suggested: the first sentence now excepts the join, and the second names OPENDJ_REPLICATION_TYPE=simple.

unhealthy and "can take". Taken as suggested.

@vharseko
vharseko requested a review from maximthomas October 9, 2026 07:32

@maximthomas maximthomas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise: The round-2 fixes were checked against the code rather than pasted, and the two departures from the suggestions are right.

  • "with the indexes as the failed rebuild left them" (chap-upgrade.adoc:372) is more accurate than the suggested "still untrusted", since a rebuild that fails before it starts leaves them trusted.
  • The pointer to /opt/opendj/data/logs/upgrade.log gives the reader the details that docker logs leaves out.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

docker docs enhancement upgrade Upgrading between versions and migrating from other directory servers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants